Papers by Francesco Maria Molfese

3 papers
Exploring Fine-Tuning for In-Context Retrieval and Efficient KV-Caching in Long-Context Language Models (2026.eacl-short)

Copied to clipboard

Challenge: Long-Context Language Models (LCLMs) can encode entire document collections, offering a strong alternative to retrieval-augmented generation (RAG).
Approach: They propose to use LCLMs to encode documents with context windows of millions of tokens to improve their performance.
Outcome: The proposed training strategies improve long-context performance and their robustness under compression techniques.
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Recent work in language modeling has led to effective SLMs with impressive performance levels across various benchmarks.
Approach: They propose a benchmark that introduces process-level evaluation for commonsense reasoning tasks.
Outcome: The proposed benchmarks show that large language models provide correct answers despite flawed reasoning processes in a substantial portion of cases.
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering (2025.findings-acl)

Copied to clipboard

Challenge: Multiple-choice question answering tasks are one of the most commonly used tasks for evaluating Large Language Models (LLMs).
Approach: They analyze whether existing answer extraction methods are aligned with human judgment and how they are influenced by answer constraints in the prompt across different domains.
Outcome: The proposed evaluation strategies can be inconsistent with human judgment, and can lead to inaccurate and misleading comparisons.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations